BMC Genomic Data
○ Springer Science and Business Media LLC
Preprints posted in the last 90 days, ranked by how well they match BMC Genomic Data's content profile, based on 13 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Fouere, C.; Costes, V.; Besnard, F.; Le Danvic, C.; Patry, C.; Fritz, S.; Boussaha, M.; Jouin, M.; Boichard, D.; Kiefer, H.; Costa Monteiro Moreira, G.; Sanchez, M.-P.
Show abstract
Background Complex traits are influenced by numerous variants, most of which have regulatory effects on gene expression that can be mediated by DNA methylation. Molecular QTL mapping is an approach that aims to dissect these effects. However, obtaining molecular phenotypes on a large scale is challenging, particularly in livestock species. In cattle, an epigenotyping array called EpiChip has recently been developed in the European RUMIGEN project. The EpiChip, which contains 43,317 CpG sites distributed all over the bovine genome, enables large-scale measurement of DNA methylation. This study aims to characterize the genetic determinism of blood DNA methylation in cows by estimating heritability and mapping cis- and trans-methylation QTLs (meQTLs). Results Whole blood samples from 4,457 genotyped Holstein cows were epigenotyped. Across all CpG sites, the heritability estimates averaged 24.6%. The local meQTL mapping at sequence-level for variable CpG sites (SD > 2.5%; n = 28,806) detected cis-meQTLs for 80.1% of the CpG sites, with sentinel SNPs located close to their associated CpGs. A two-step analysis was also conducted to identify long-range associations, with a particular focus on trans-meQTL hotspots. First, we identified CpG-SNP trans-associations using medium-density genotypes (50k SNPs) that revealed 31,846 SNPs with significant effects on 1 to 530 trans-CpG sites. Then, regions associated with at least 34 independent trans-CpGs were retained defining 31 hotpots. For each hotspot, a local sequence-level GWAS was conducted using the first principal component derived from the associated trans-CpGs. Out of the 31 detected hotspots, three were located close to transcription factor genes (RUNX1, NFIC and FOXA3) for which the associated trans-CpGs were enriched for the corresponding binding motif. Two other hotspots were located within KDM5A and KDM5B, and their corresponding trans-CpGs were strongly overrepresented in H3K4me3 narrow peaks in blood as well as in other tissues. Conclusions By identifying functional candidate genes associated with blood DNA methylation in cattle, these findings provide new insights into the regulatory architecture of DNA methylation in mammals, highlighting the value of large-scale molecular data from livestock populations.
Lanigan, S.; Derks, M. F.; Johansson, A. M.; Johnsson, M.
Show abstract
Variant intolerance methods score the essentiality of genes based on large datasets of genetic variants and have been used in population genomics of humans and model organisms. In this paper, we estimated Residual Variation Intolerance Scores for protein-coding genes and predicted protein domains in cattle. In agreement with results from other species, the most variant-tolerant genes and domains included genes related to olfaction and adaptive immunity, whereas the least-variant tolerant genes and domains included genes involved in fundamental cellular processes. There was a moderate positive correlation with estimates from orthologous human genes. We provide estimates of variant intolerance for cattle may be useful for genomic analyses of deleterious variants and population genomics in cattle.
Sundelin, H.; Jacobsson, B.; Ytterberg, K.; Sole-Navais, P.; Juodakis, J.
Show abstract
The leading cause of mortality and morbidity in children under the age of 5 is preterm birth. The timing of birth is influenced by both genetic and environmental factors, but the underlying mechanisms remain poorly understood, making its prediction difficult. In this study, we investigated the potential of using machine learning models to predict preterm birth based on genetic data from the Norwegian Mother, Father and Child Cohort Study (MoBa). We trained and evaluated several classification algorithms on individual-level genetic data from over 15,000 mothers and children. Our results indicate that the predictive capacity of maternal gestational duration-associated loci for preterm birth is limited, with the highest AUC values around 0.57. Additionally, incorporating more SNPs within the associated loci did not improve prediction performance. As expected, the contribution of the maternal genome to preterm birth prediction was found to be larger than that of the fetal genome. Overall, our findings suggest that while genetic testing provides some information about an individual's risk for preterm birth, further research incorporating additional factors is necessary to enhance predictability.
Shi, L.; Ojemakinde, O. T.; Hart, C. M.
Show abstract
Chromosomes in Drosophila melanogaster are organized into distinct topologically associated domains delimited by boundaries bound by insulator proteins. The insulator protein BEAF (Boundary Element-Associated Factor of 32kDa) plays roles in both chromatin organization and transcriptional regulation, yet its precise molecular mechanisms remain elusive. To examine the role of BEAF in insulator function we used yeast 2-hybrid assays, pull-down assays with bacterially expressed proteins, and bimolecular fluorescence complementation assays in S2 cells to characterize interactions between BEAF and three co-insulator proteins: CP190, Pzg, and Chro. Our studies pinpoint minimal regions of CP190, Pzg, and Chro that directly interact with BEAF, as well as parts of BEAF that are crucial for its interactions with these co-insulator proteins. Functional analyses in transfected S2 cells revealed distinct regulatory roles for BEAF in association with CP190, Pzg, and Chro. CP190 showed a weak interaction with BEAF, and CP190 bound 2.3 kb upstream of promoter-proximal BEAF could not loop out the intervening DNA to effectively communicate with BEAF for luciferase reporter gene activation. In contrast, more robust interactions were observed between BEAF and both Pzg and Chro. Both could also effectively interact with BEAF from a distance to activate the reporter gene. It is likely that the role of CP190 in long-range insulator interactions is mediated by insulator proteins other than BEAF, although BEAF could help stabilize interactions. On the other hand, we propose that BEAF can directly collaborate with Pzg and Chro to mediate long-range chromatin interactions.
Magateshvaren Saras, M. A.; Ahmad, S.; Smith, R.; Mitra, M. K.; Tyagi, S.
Show abstract
The early onset of labour increases mortality and developmental risks for a human newborn. Key genes in human labour have been investigated using multiple modalities, but their regulation by non-coding RNA (e.g. lncRNA and miRNA) remains incomplete. This study explores the three-way relationship between labour-associated transcription factors (TFs), miRNA and lncRNA suggested by the competing endogenous RNA (ceRNA) hypothesis, to understand the underlying regulatory framework. Experimentally validated miRNA-lncRNA interactions are modelled using five distinct machine learning (ML) architectures to predict 20469 labour-linked miRNA-lncRNA interactions. Known mRNA-ncRNA interactions from databases were included to construct a tripartite network, and a subset of 9989 labour-linked network motifs containing TFs were isolated and analysed. Gene enrichment of nodes in TF-lncRNA-miRNA network, as well as validation from public myometrial datasets indicate high significance in contractile pathways including immune signalling. Experimentally unconfirmed tripartite network motifs have been found, and we elaborate on their potential regulation in labour using 8 TF-lncRNA-miRNA network motifs. A unified ncRNA-TF regulatory atlas in labour has been synthesized, and a complete summary of the tripartite network motifs can be accessed and visualised using the user-friendly, public database.
Shrestha, A. M. S.; Manlapaz, J. P.
Show abstract
Numerous genome-wide association studies in rice have identified loci associated with diverse agronomic traits. However, interpreting the regulatory and functional significance of these loci remains challenging because each locus often contains multiple variants in linkage disequilibrium, many of which lie in non-coding regions. Here, we present a method for prioritizing non-coding variants within an associated locus by integrating chromatin feature information predicted by a pretrained DNA language model fine-tuned on rice ChIP-seq and ATAC-seq datasets. We demonstrate the utility of our method through three case studies. In a post-GWAS analysis of heat tolerance, prioritized variants overlapped promoters of candidate genes previously identified through integrated GWAS and transcriptomic analyses, providing independent support for their potential regulatory roles. For the high-yield gene DEP1, promoter variant prioritization combined with in silico saturation mutagenesis identified a localized regulatory region enriched for high-impact mutations overlapping predicted transcription factor binding sites. For the drought-associated gene OsHAK1, the highest-ranked variant was predicted to be associated with chromatin features in a manner consistent with the reported co-occurrence of H2Bub with H3K4 methylation marks in plants. Overall, these results demonstrate the utility of our method for functionally informative prioritization of non-coding variants, facilitating the interpretation of GWAS loci and the identification of candidate regulatory variants in rice.
Zhang, Z.; Xu, Y.
Show abstract
Language genes can be tentatively considered as a subset of cognitive genes, although they are often discussed separately. During the evolution of SNVs (single nucleotide variations) in cognition-related genes, do language genes and cognitive genes exhibit significantly different intensities of change at several key evolutionary moments--namely, the inflection points or derivative peak positions of similarity curves drawn from multi-SNV locus bases across samples? In this study, nine distance/similarity metrics (Bray-Curtis, Cosine, Pearson, Spearman, Hamming, Jaccard, Matching, Kulczynski, and Gower) were employed to analyze 413 samples from 11 taxonomic groups, targeting SNV loci in language/cognition-related genes (13,415 effective loci, approximately 400 loci per gene), with pp6 (Homo_sapiens.GRCh38) as the reference. For each method, sample similarities (defined as 1/(1+distance)) were independently sorted in ascending order to generate raw similarity scatterplots. Due to the large sample size and representativeness, the scatter density on the similarity curves was high, and no smoothing was applied. Derivative values were calculated from adjacent similarity differences to identify peaks of evolutionary rate change (top 10 peaks per method). Combined with functional annotations of 33 language/cognition-related genes, we quantified the difference scores and occurrence frequencies of the two gene categories at the peak positions. The results indicate that cognitive-related genes exhibit slightly higher occurrence frequencies in peak windows and higher average difference scores per gene than language genes. Comparative analysis of SNVs at the peak samples and their left-side windows revealed that at positions 381-382, all nine methods shared three intersecting mutation loci, involving language genes (NFXL1, SRGAP2, SRGAP2C); at positions 355-356, there was one intersecting mutation locus, involving a language gene (SRGAP2). This suggests that certain mutations in language genes may have played a distinctive role at critical junctures in the evolution of cognitive abilities.
Mandic, K.; Hrsak, D.; Uljanic, F.; Lenhard, B.; Baresic, A.
Show abstract
Genome-wide association studies (GWAS) are the key tools for the discovery of associations between single nucleotide polymorphisms (SNPs) and phenotypic traits and have been successfully applied to many diseases and disorders. However, a great challenge is to find the gene affected by the non-coding fraction of SNPs, especially if the gene is distal in terms of genomic distance. In this study, we present a novel approach, named targPred, which utilises genomic regulatory blocks (GRBs) for inference of a connection between a certain SNP/locus and the target gene located in the same GRB, in a more robust and generalisable manner. We identified that many disease traits such as cancer and psychiatric disease have a propensity for long-range regulation. Furthermore, we showcased a childhood obesity locus which is connected to the distal BDNF gene. Finally, we propose a new web-based service based on enhancer-promoter association, to facilitate finding the causal genes for a wide array of traits and conditions.
Matarage Don, N. N. J.; Biswas, S. B.; Biswas-Fiss, E. E.
Show abstract
Pathogenic mutations in the ABCA4 gene cause several inherited retinal diseases, particularly Stargardt disease (STGD1). However, many missense variants remain classified as variants of uncertain significance (VUS) due to inconclusive evidence regarding their pathogenic impact. The missense VUS span across all the domains of ABCA4, with the majority found in the larger extracellular domains (ECDs). The largest uncharacterized region of ABCA4 is located in ECD1, where limited structural information and inconsistent computational predictions hinder clinical interpretation of missense VUS in this region. Here, we integrated in silico analysis with in vitro functional assays to evaluate the pathogenicity of VUS in this region and improve their diagnostic classification. Missense VUS in the ECD1 uncharacterized region were curated from ClinVar. Six multiallelic sites were identified in the uncharacterized region and 13 missense VUS on these multiallelic sites were characterized using the integrated analysis. In the in silico platform, the pathogenicity of the VUS were predicted using multiple algorithms, and the structural effects of the variants were analyzed compared to the wild type. Recombinant variants were expressed in virus-like particles (VLPs), and protein expression, membrane localization, and ATPase activity were quantified relative to wild type to identify potential disease-causing variants. From the integrated analysis, variants with pronounced structural destabilization, impaired membrane trafficking, and reduced or absent N-retinylidene-phosphatidylethanolamine (NRPE) substrate stimulated ATPase activities were identified as potentially deleterious. Notably, VUS at p.H193P and p.I214N showed loss of function, with p.I214N reflecting selectively impaired membrane targeting and p.H193P reflecting combined expression and trafficking defects. Additionally, NRPE-stimulated ATPase activities were impaired in VUS, p.V195L, p.V195I, p.D197H, p.I214F and p.N269S. Overall structural destabilization interfered with the NRPE-stimulated ATPase activities of p.N269S, while the lack of NRPE-stimulated ATPase activities of p.D197H, p.V195L, p.V195I and p.I214F are thought to be due to impaired NRPE interactions with ABCA4. All the VUS at p.R140, p.H193Y, p.D197N and p.N269H showed both the basal and NRPE-stimulated ATPase activities but less than that of the wild type, displaying a mild functional deficit. Together, these findings demonstrated that certain VUS within the unresolved ECD1 region disrupt ABCA4 stability and function, supporting their contribution to disease pathogenesis. This integrative approach highlights key residues likely to be pathogenic and advances the interpretation of VUS in inherited retinal disorders.
Zhang, Z.; Xu, Y.
Show abstract
This study aims to quantify the genetic similarity of different species (from fish to humans) to the human reference genome (pp6, Homo sapiens.GRCh38) based on the allele presence/absence patterns of 33 language/cognition related gene SNV loci, identify key breakpoints during evolution, and evaluate the enrichment of language and cognition genes at these breakpoints. We designed a similarity calculation method relying on binary features (four columns for A/T/C/G), adopted five difference/distance measures (Sorensen, Rogers, Nei, Reynolds, and Hellinger), and converted them into similarity values (1/(1+distance)). For each method, samples were independently ranked, the first derivative of similarity was computed, and the top 12 peaks were selected as candidate breakpoints. Results show that the similarity curves from the five methods are highly consistent (correlation coefficients >0.9), with major peaks concentrated at positions 355, 363, 381, 382, 390, 400, etc., where the corresponding samples are predominantly ancient hominins and primates. Furthermore, we defined 13 peak groups (starting positions 355-401). For each peak within a group, pairwise SNV differences between the peak apex sample and its immediate left neighbor were compared, and the intersection F_INTERSECTION (shared differential loci) was obtained. For each F_INTERSECTION, we calculated the proportions of language genes and cognition genes. In addition, we computed the differential sets between adjacent groups' F_INTERSECTION to trace the gradual emergence of new loci. In F_INTERSECTION, language genes accounted for an average of 59.5%, and cognition genes for an average of 62.9%. The proportion of language genes reached a peak at position 383 (61.2%), while cognition genes peaked at position 386 (64.9%). High frequency peak samples include c25, c27, and ja2, suggesting that language cognition genes may have undergone independent intensification during Eurasian evolution. Differential analysis between adjacent F_INTERSECTION revealed a stepwise acquisition of new loci from position 355 to 401, with three bursts of newly added loci along the entire evolutionary axis. This study provides a quantitative framework based on similarity curves, offers a novel molecular perspective for understanding the evolution of language and cognitive abilities, and highlights the potential importance of East Asian archaic hominins in the evolution of language cognition genes.
Arya, A.; Datta, B.
Show abstract
Symmetry elements in nucleic acids are most strongly correlated with sites of biological function; however, their relevance to non-canonical structures remains underexplored. In this study, we demonstrate the presence and significance of trinucleotide symmetry elements within G-quadruplex (G4) motifs. Our central hypothesis is that the intra-strand mirror symmetry of trinucleotides has been evolutionarily selected to facilitate G4 formation builds on the established sequence-structure association of G-quadruplexes and the natural symmetry law governing nucleotide insertion during genome evolution. Using a conserved G4 motif in the first exon of the MTOR gene as a model, we showed remarkable trinucleotide symmetry preservation across primates and broader mammals, with functional G4 regions displaying locally elevated symmetry relative to the codon-biased exonic background. Analysis of experimentally validated oncogenic G4s, including c-MYC, BCL2, VEGF, and KRAS, revealed that mirror and reverse complement symmetries converge around biologically important G4s. To quantify this feature, we formulated two complementary descriptors: the mirror symmetry index (MSI) and its non-palindromic variant (nMSI). Across 14 oncogene-promoter wild-type G4s, the majority scored MSI [≥] 0.80 (mean 0.884), with only the loop-rich ATG7, BCR, and MDM2 motifs falling below this value, and the KRAS promoter G4 reached individual significance against its mononucleotide-preserving null distribution (p = 0.042). Most decisively, each wild-type G4 scored higher on MSI than its experimentally confirmed G4-abolished mutant in 12 of 14 paired comparisons (sign test, p = 0.0065; mean {Delta}MSI = +0.089, mean {Delta}nMSI = +0.192); the two reversals (BCL2 and HIF-1) are attributable to scrambled mutant controls that introduce more balanced trinucleotide compositions rather than to failure of the index. The directional trend was reproduced across three independently published datasets, with nMSI [≥] 0.50 separating G4-forming from non-G4 sequences at 77.8% sensitivity and 100% specificity, although the collective per-sequence signal from mononucleotide-preserving shuffles remained a non-significant trend (Stouffer combined Z = 1.197, p = 0.116). This first report of trinucleotide symmetry in G4 motifs posits that coordinated nucleotide insertion and quadruplet maintenance act as an evolutionary forcing mechanism that pre-organizes single strands for G4 folding.
Kandasamy, R.; Gurung, M.; Shrestha, S.; Bibi, S.; Thorson, S.; Carter, M.; O'Connor, D.; Murdoch, D. R.; Kelly, D. F.; Shrestha, S.; Levin, M.; Pollard, A. J.
Show abstract
Background Pneumococcal disease is a leading cause of paediatric pneumonia and meningitis. Pneumococcal colonisation is the fundamental step to pneumococcal disease causation. We aimed to identify genetic loci associated with pneumococcal colonisation amongst children. Methods We conducted a genome-wide association study on 2111 Nepalese children, comprising 1346 cases carrying pneumococcus and 765 controls. We tested 8.1 million imputed variants using logistic regression and ten principal components as covariates. Fine mapping and functional evidence were used to identify suspected causal variants and related genes of interest. Findings A cluster of 22 variants of genome-wide significance (p<5x10-8) were identified on chromosome 12q21.31, eight of which were within PPFIA2. Fine mapping of this region identified 5 variants within 0.1 Mb of the 5-prime region of PPFIA2 all of which are significant eQTLs for PPFIA2. We further describe three loci (10q23.31, 12q23.1, and 20p11.21) which had variants with highly suggestive associations (p<5x10-7)with pneumococcal carriage. Interpretation Our study demonstrate human susceptibility to pneumococcal carriage to be polygenic with genetic variations which regulate PPFIA2 expression playing a key role in the ability for pneumococcus to colonise children. Targeting these genetic factors and the associated pathways are a means for preventing pneumococcal disease. Funding This study was supported by funding from Gavi - the vaccine alliance, the European Unions Horizon 2020 research and innovation program under grant agreement number 668303 (PERFORM), and a Robert Austrian Research Award.
Mazgaj, R.; Kołpa, A.; Esmaeeli, M.; Pełczynska, J.; Galea, D.; Gawor, J. J.; Malinowska, A.; Szczypiorowska, A.; Kehl-Fie, T.; Waldron, K. J.
Show abstract
Background: Biochemical, biophysical and structural characterisation of isozymes from the ubiquitous family of iron- or manganese-dependent superoxide dismutases (SodFMs) requires the purification of high-quality preparations of recombinant enzymes. Determination of their key biochemical parameter, their catalytic metal-preference, requires the comparison of the catalytic turnover of samples loaded exclusively with iron versus samples loaded exclusively with manganese. Both of these aims are inhibited by the potential contamination of recombinant preparations of SodFMs, prepared by heterologous overexpression inside Escherichia coli cells, by even low levels of endogenous SodFMs from the host, both of which show very high turnover with either manganese (E. coli MnSOD) or iron (FeSOD). To overcome this problem, we created a strain of E. coli lacking the endogenous SodFMs. Here, we characterised this E. coli BL21 (DE3) {Delta}sodA{Delta}sodB strain, determining the physiological effects of SodFM deletion and demonstrating its utility for producing recombinant SodFMs for in vitro characterisation and use. Results: Genomic analysis verified the targeted gene deletions, without off-target effects. Growth, expression, elemental analysis, and proteomic data confirmed a lack of physiological defects of the strain except for a known inability to grow on glucose, which is overcome by heterologous SodFM expression. We demonstrate the utility of the strain for the efficient production of diverse recombinant SodFMs, including highly divergent, understudied isozymes, including the ability to precisely control the metal-loading of the heterologously expressed protein. Conclusions: The E. coli strain described herein is a useful microbial cell factory for production of recombinant SodFMs, which should find widespread utility as expression host of choice, enabling more efficient production of protein for studies of the biochemical, biophysical and structural properties of this remarkable family of metalloenzymes.
de Leeuw, V. C.; Maitre, L.; van Oostrom, C. T.; Renard-Dausset, E.; Anguita, A.; Chatzi, L.; Coen, M.; Grazuleviciene, R.; Heude, B.; Ibarluzea, J.; Julvez, J.; Keun, H. C.; Piersma, A. H.; Maria, L. S.; Marquez, S.; Ruiz-Rivera, M.; Subiza-Perez, M.; Brantsaeter, A. L.; Toledano, M. B.; Vrijheid, M.; Wright, J.; Hessel, E. V.; Hoyles, L.; McArthur, S.
Show abstract
Interest in microbiota-host co-metabolism and the effects of its derived co-metabolites on biological processes is increasing rapidly. In addition to their demonstrated associations with mammalian metabolic health and cognition, microbiota-host co-metabolites (MHCMs) represent lifelong contributors to the endogenous exposome. We have previously shown the MHCM trimethylamine N-oxide (TMAO) to exert beneficial effects on murine blood-brain barrier integrity and cognition. Here we investigated whether these positive neural effects of TMAO extended to humans, analysing how TMAO exposure associates with neurodevelopmental outcomes in children and whether an in vitro human neuronal-astrocyte co-culture could contribute to further investigation of the underlying mechanism(s) and neuronal processes related to these associations. In a cohort study of childhood mental health (N=1,203), TMAO was associated with fewer internalising problems, while its precursor microbial metabolite trimethylamine was associated with more behavioural problems in both the cross-sectional and an independent longitudinal study from 1 to 15 years of age (N=630-820). Given prior associations between TMAO exposure and exposure to the environmental pollutants mercury and arsenic, we investigated how the effects of TMAO interacted with these known neurotoxicants. TMAO had a protective effect, modifying the relationship between arsenic exposure and poorer neurodevelopmental outcomes. Furthermore, TMAO activated synaptogenesis-related gene expression and was functionally protective against the negative effects of mercury in our in vitro model. Together, our findings emphasise the importance of interdisciplinary approaches to evaluate associations and potential pathways of MHCMs (endogenous) and environmental (exogenous) metabolites on neurodevelopment in exposome studies.
Creasey, L. D.; Tauber, E.
Show abstract
RNA methylation at N6-adenosine (m6A) predominantly occurs within RRACH motifs, yet the forces shaping these motifs in coding regions remain unclear. Here we show that the synonymous codon combinations able to form or disrupt RRACH sites are used non-randomly across mammals. Using 13,491 protein-coding genes from 261 species, we identified genes significantly enriched or depleted in RRACH motifs, a pattern consistent with gene-specific selection for or against m6A potential. Genes enriched in RRACH sites were linked to ubiquitin-like conjugation and cell cycle regulation, whereas transmembrane and HOX genes were RRACH-poor, likely reflecting sequence incompatibility with CpG dinucleotides. Cross-species comparison with Caenorhabditis elegans, which lacks mRNA m6A methylation, revealed reciprocal RRACH frequencies, as expected if these motifs are under selection in m6A-competent genomes but evolve without this constraint otherwise. At the codon level, specific amino acid pairs, particularly threonine-ending dyads, were biased toward RRACH-forming codons while others were depleted, indicating that synonymous codon choice is skewed for and against motif formation. RRACH motifs were also non-randomly distributed along coding sequences, depleted near start codons and enriched toward the 3' end, consistent with known m6A profiles. Finally, analysis of cancer mutations revealed tissue-specific gain and loss of RRACH sites, reflecting context-dependent remodeling of methylation potential. Together, these results show that synonymous codon usage is systematically biased for and against m6A RRACH motifs, pointing to an evolutionary coupling between the genetic code and the epitranscriptomic landscape.
Singer, A. L.; January, E. E.; Zess, E. K.; Antonakos, A. J. N.; Begemann, M. B.
Show abstract
Cas12a2 CRISPR nucleases, including SuCas12a2, have been shown to have extensive collateral activity towards RNA, ssDNA, and dsDNA. This collateral activity results in targeted cell elimination and has applications across biotechnology, agriculture, and human health. We explored the natural genetic diversity of Cas12a2 nucleases and characterized nine novel orthologs in a DNA damage kinetic assay in E. coli. Three new Cas12a2 orthologs (RsCas12a2, SdCas12a2, and HmCas12a2) were shown to have high collateral activity towards DNA. These nucleases are highly divergent from SuCas12a2, have conserved core RuvC catalytic residues, and have sequence diversity in the previously reported aromatic clamp residues required for nucleic acid positioning in the active site. We defined PFS preferences and mismatch tolerance for each high-activity Cas12a2 nuclease, expanding the available Cas12a2 toolbox, and discovered functional differences with obvious impacts on downstream applications.
zhang, Y.; LI, K.
Show abstract
Background: Quantitative assessment of immune function is essential for clinical and health decisions in oncology, post-surgical management, and autoimmune diseases. Existing methods are either too simplistic (single indicators) or too complex and costly for routine use. A standardized, easy-to-operate tool based on routine laboratory parameters is needed for both clinical and health checkup settings. Methods: We propose the Immune Index (II), integrating 9 routine laboratory parameters across three dimensions: humoral immunity (IgG, complement C3, C4), cellular immunity (CD4+ T cells, CD8+ T cells, CD4+/CD8+ ratio), and inflammatory response (CRP, IL-6, systemic immune-inflammation index [SII]). Indicators were normalized using min-max normalization to a 0-100 scale and aggregated with fixed weights (humoral 30%, cellular 40%, inflammatory 30%). The II score ranges from 0 to 100, with a healthy reference range of 50-80. Results: A four-tier grading system was established: >=80 (immune overactivation), 50-80 (immune homeostasis), 35-50 (mild immune suppression), <35 (severe immune deficiency). Validation using 209 cases from published literature showed an AUC of 0.924 (95% CI: 0.87- 0.97) for distinguishing normal from abnormal immune status, with an optimal cutoff of 47.8 (sensitivity 84.8%, specificity 85.9%). II scores were 56.7+/-8.6 (healthy), 43.5+/-8.0 (immunodeficient), and 33.6+/-6.5 (autoimmune), with P<0.001 between all groups. The calculation requires only two steps and can be implemented in Excel or LIS. II can serve as an immune dimension supplement for personal health checkups. Conclusion: The Immune Index provides a simple, standardized, and low-cost tool for quantitative immune function assessment. The fixed-weight design ensures cross-institutional comparability, making it suitable for outpatient clinics, health checkup centers, and primary care settings. Keywords: Immune index; immune function; quantitative assessment; routine laboratory parameters; composite score; min-max normalization
Zhang, B.
Show abstract
Toxicants in the environment can significantly impact physiology. Environmental chemical exposures during early developmental stages disturb normal embryonic development and programming, and dramatically impact long-term health as individuals age. Female and male animals show distinct phenotypes when responding to a given chemical exposure. Here, through the TaRGET II (Toxicant Exposures and Responses by Genomic and Epigenomic Regulators of Transcription) consortium, we systematically explored sex-specific transcriptomic and epigenomic alterations in response to various toxicants, including arsenic (As), lead (Pb), tributyltin (TBT), bisphenol A (BPA), di(2-ethylhexyl) phthalate (DEHP), dioxin (TCDD), and fine particulate matter (PM2.5), across three time points in mice exposed two weeks prior to conception through gestation and lactation. After being exposed to toxicants during the embryonic and early postnatal developmental stages, 1,025 omics datasets were generated from the liver and analyzed across three mouse life stages. We discovered a significant sex-biased molecular response to distinct exposures in the liver at both the transcriptomic and epigenetic levels, showing dynamic changes across mouse development and aging. The perturbed pathways and transcription factors in response to different chemical exposures in both sexes were further evaluated to measure the sex-specific impact of each toxic exposure in the liver. Overall, this study presents the most detailed investigation of sex-specific molecular signatures under the influence of developmental exposures to toxic substances.
Layman, C. E.; Morrow, D.; Wheeler, K.; Caron, T. J.; Davis, B. A.; Bergstrom, P.; Vigh-Conrad, K.; Anderson, T. J.; McElfresh, G. W.; Sterner, K. N.; Sadoughi, B.; Snyder-Mackler, N.; Hansen, S. G.; Bimber, B. N.; Lancioni, C.; Carbone, L.; Okhovat, M.
Show abstract
Wildfire smoke is an escalating global public health threat exposing millions of people, including children, to hazardous air pollution each year. Although wildfire smoke toxicants have been linked to a range of adverse health outcomes, including immune dysregulation, the long-term consequences of real-world pediatric wildfire smoke exposure on health and development remain largely unknown. To investigate the persistent effects of early-life exposure on immune health, here we leveraged a cohort of rhesus macaques that experienced nine consecutive days of hazardous wildfire smoke exposure in infancy during the 2020 Oregon Labor Day wildfires. By integrating ex vivo immune stimulations, multiplex cytokine profiling, single-cell transcriptomics, and genome-wide DNA methylation profiling, we identified persistent immunological consequences across molecular and functional levels. We found that a single severe postnatal exposure, in the first three months of life, was associated with persistent change in the innate immune response, including reduced pro-inflammatory cytokine response to a bacterial endotoxin, with subtle but consistent transcriptional changes in myeloid cells, particularly among males. Wildfire smoke exposure was also associated with changes in proportion of B and T/NK cells, and within the T/NK cell compartment, exposed animals exhibited an expansion of cytotoxic cells. Consistent with this, CD8+ T cells displayed extensive transcriptional remodeling and shifted toward more differentiated effector states, with the greatest differentiation observed in animals exposed at the youngest ages. Genome-wide DNA methylation profiling identified smoke-associated methylation changes consistent with acceleration of epigenetic aging, as well as persistent epigenetic alterations impacting genes involved in oxidative stress responses, innate immunity, T cell differentiation, and hematopoiesis. These findings demonstrate that a single severe wildfire smoke exposure during a critical developmental window is associated with extensive immune and epigenetic remodeling that persist years after exposure, providing new insight into the long-term biological consequences of early-life wildfire smoke exposure.
Jesudasan, R.;Mukhoti, A.;Chaturvedi, A.;Tiwari, S.;Mishra, K.;Pranatharthi, A.;Praveena, N.;Alex, J.;Karunanithi, S.;Kumar, A.;Reddy, H.
Show abstract
BackgroundHeterochromatic long arm of mouse Y chromosome harbors the multicopy species-specific sequences Ssty, Sly, Asty and Orly that are transcribed in testis and have known functions in male fertility. Of these Ssty and Sly encode proteins - yet all the transcripts are not translated. To investigate the roles of these Y-heterochromatic transcripts further, we analyzed them. MethodsMice with 2/3rd deletion of the Y-chromosome (XYRIIIqdel) and its wild type (XYRIII) were used in this study. Bioinformatic approaches, small RNA northern blots, Electrophoretic Mobility Shift Assays, Luciferase reporter assays, dPCR analysis, RT-qPCR assays and western blotting techniques were used to identify piRNAs that regulate autosomal genes. ResultsWe demonstrate that the multicopy gene families from mouse Y-long arm generate piRNAs predominantly in testis. We observed sequences homologous to these piRNAs in the UTRs of a few autosomal genes, which are differentially expressed in the sperms of XYRIIIqdel mice. Furthermore, the Endogenous Retrovirus Element (ERV) LTR, found in the Orly1 transcript identified piRNAs in the database, showed homology to UTRs and associated genomic regions of a few autosomal genes. Orly1 showed a reduction in genomic copy number by digital PCR in XYRIIIqdel mice. One of the four autosomal genes containing the ERV segment in their UTRs, showed a differential testicular protein expression in the mutant mice. ConclusionsThus, we further elucidate that different classes of repeats from Y-chromosome regulate autosomal gene expression via piRNAs. Besides, this study also identified novel roles for a Y-derived ERV in autosomal gene regulation in testis.